Repository navigation
Conversation
Previously to see the same error in docs build as you would get in CI, you would have to run both the docs build and the spellchecking build. The issue was that the spellcheck (using sphinx-contrib's spelling) was defined as a new builder, meaning you could either build HTML, or you can check spelling. On CI, this didn't result in any extra wall-clock time as the jobs were parallelized (but this change does remove a job from the overall matrix), but locally that was almost always not the case, so this would mean full build to view the HTML, and spell checking build to look for errors, resulting in it taking twice as long. The upstream spellchecker was also slower than it should be mainly because its importable-module filter calls find_spec on every distinct identifier-like word before consulting the dictionary. Here it runs only for words that are still misspelled. Word lists, filters and `:spelling:ignore:` behave as before. This also handles errors a lot more nicely (hence all the code deletions): - They are written natively to a JSON file, not scraped from a stream - They also show up as GH annotations (Where we can attribute it to a real file) to make fixing them easier on PRs. Please note: this is lightly breaking for Airflow contributors as `breeze build-docs --docs-only` and `--spellcheck-only` are gone -- every build checks spelling now as it is essentially free (on the order of milliseconds per module).
ashb
requested review from
amoghrajesh,
bugraoz93,
choo121600,
ephraimbuddy,
gopidesupavan,
hussein-awala,
jason810496,
jedcunningham,
jscheffl,
potiuk and
vatsrahul1001
as code owners
October 10, 2026 21:24
ashb
force-pushed
the
single-pass-spelling
branch
2 times, most recently
from
October 11, 2026 09:19
b6c14c7 to
733b47b
Compare
Member
|
WHOA !! Fantastic !! Thanks @ashb !!!! |
Member
|
And you can see the errors in CI !! . We should have thought about it way before !!! |
Member
Author
I honestly thought this would be way more difficult than it was. I had a lot of time waiting keeping an eye on the Task Loops stack (now thankfully landed!) and yeah, this turned out to be much easier. The one caveat to the error annotations: These do not ("cannot" right now) show up for generated API docs, as we don't have a way to to map api docs to source file it came from |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Previously to see the same error in docs build as you would get in CI,
you would have to run both the docs build and the spellchecking build.
The issue was that the spellcheck (using sphinx-contrib's spelling) was
defined as a new builder, meaning you could either build HTML, or you
can check spelling. On CI, this didn't result in any extra wall-clock
time as the jobs were parallelized (but this change does remove a job
from the overall matrix), but locally that was almost always not the
case, so this would mean full build to view the HTML, and spell checking
build to look for errors, resulting in it taking twice as long.
The upstream spellchecker was also slower than it should be mainly
because its importable-module filter calls find_spec on every distinct
identifier-like word before consulting the dictionary. Here it runs only
for words that are still misspelled. Word lists, filters and
:spelling:ignore:behave as before.This also handles errors a lot more nicely (hence all the code deletions):
real file) to make fixing them easier on PRs.
Please note: this is lightly breaking for Airflow contributors as
breeze build-docs --docs-onlyand--spellcheck-onlyare gone --every build checks spelling now as it is essentially free (on the order
of milliseconds per module).
Was generative AI tooling used to co-author this PR?
{pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.